Back

Medical Image Analysis

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Medical Image Analysis's content profile, based on 35 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
NeuroMesh: A Bottleneck Topology Controller for Missing-Modality Brain Tumor Segmentation - A Mechanistic Pilot Study on BraTS

Kamalakannan, N. K.; Kamalakannan, J.

2026-08-27 bioinformatics 10.64898/2026.08.23.746542 medRxiv
Top 0.1%
34.2%
Show abstract

Deep segmentation networks can degrade sharply when an expected MRI sequence is unavailable at inference. We present NeuroMesh, a bottleneck controller that combines a gated recurrent unit (GRU) with a graphconvolutional edge-activation mask, designed to adapt a U-Net-style segmentation backbone to missing input. We evaluate NeuroMesh in a pilot study using a 30-patient subset of the BraTS 2020 benchmark (22 training, 4 validation, and 4 held-out test patients) under a prespecified frozentest protocol. On the frozen test set, NeuroMesh has higher tumor-core and enhancing-tumor Dice than a plain U-Net in most evaluated missing-modality conditions, but wholetumor Dice falls from 0.596 to 0.108 when FLAIR is missing, compared with 0.604 to 0.545 for the plain U-Net. Direct analysis of the predicted edge-activation mask shows negligible change across modality-availability conditions. A parameter-light static-gating control reproduces the FLAIR failure mode without recurrence, a failure-signal input, or graph-structured machinery. These results do not support the intended interpretation that the trained controller performs input-conditional topology rewiring at the scale of this pilot. Instead, they expose a discrepancy between architectural intent and realized behavior and identify a specific missing-modality failure mode that warrants further investigation. Given the small validation and test sets, the findings are descriptive and do not establish clinical or population-level generalization.

2
Reconstructing synthetic hearts from ECG using flow matching

Zheng, J.; Kalaie, S.; Ma, Q.; Meng, Q.; Rjoob, K.; Gifani, P.; Hu, L.; Babazade, N.; Coriano, M.; Zhong, W.; Vafaeezadeh, M.; Tahasildar, S.; Vadgama, N.; Senevirathne, D. S.; Santhirasekaram, A.; McGurk, K. A.; Curran, L.; He, Y.; Chen, L.; Mo, Y.; Huang, L.; Qiao, M.; Huang, Y.; Bai, W.; O'Regan, D. P.

2026-09-04 cardiovascular medicine 10.64898/2026.09.01.26360987 medRxiv
Top 0.1%
33.9%
Show abstract

Cardiac imaging enables quantitative assessment of cardiac structure and function but remains constrained by cost, infrastructure and specialist expertise. In contrast, electrocardiogram (ECG) is widely accessible yet underexploited, despite encoding latent information about cardiac physiology. Here we introduce visionECG, a conditional flow matching framework that learns a probabilistic mapping between two biological distributions - the space of cardiac electrical signals and the space of cardiac geometries. Using 71,132 paired ECG and cardiac mesh sequence datasets from the UK Biobank, with external assessment in 5,000 patients with ECG-echocardiogram pairs, the model reconstructs quantitatively accurate spatiotemporal representations of the left ventricle using ECG inputs and basic demographic information alone. These reconstructions enable discrimination of structural abnormalities and disease labels, provide visualisations of functional abnormalities, and support flexible quantification of both global and regional parameters. By reframing the ECG as a generative source of patient-specific left ventricular geometry and motion, this work establishes a scalable framework for translating low-dimensional signals into high-dimensional, physiologically grounded structured representations.

3
A Vision-Language Model for Coronary Angiography Interpretation and Clinical Decision Support

Li, Z.; Sun, Y.; Jiang, C.; Pan, T.; Zhou, Y.; Wang, C.; Pan, L.; Zhang, X.; Yang, Z.; Yu, Z.; Xiao, Z.; Chen, J.; Huang, Y.; Sun, R.; Gan, Y.; Li, X.; Zhang, B.; Zhang, Z.; Wang, X.; Han, L.; Qi, Y.; Cheng, Y.; Liang, Y.; Ge, J.

2026-08-12 cardiovascular medicine 10.64898/2026.08.11.26360095 medRxiv
Top 0.1%
13.2%
Show abstract

BACKGROUND: Coronary angiography remains the reference standard for diagnosing coronary artery disease and guiding revascularization, yet its interpretation requires expert integration of multi-view anatomy, lesion morphology and procedural context. Existing artificial intelligence approaches are largely task-specific, annotation-dependent and limited in capturing the semantic relationship between angiographic findings and interventional decision-making. Whether large-scale vision-language pretraining can enable transferable foundation-model representations for invasive coronary imaging remains unknown. METHODS We developed CAG-MIND, a domain-specific vision-language foundation model for coronary angiography, using 135,475 CAG examinations paired with procedural reports, comprising 812,850 angiographic videos from Zhongshan Hospital and Shanghai Geriatric Medical Center. Each case consisted of standardized six-view angiographic acquisitions paired with structured procedural semantics extracted from routine reports using a large language model-assisted pipeline. The model was pretrained by aligning multi-view angiographic representations with report-derived semantic embeddings through bidirectional contrastive learning. Performance was evaluated under zero-shot and supervised fine-tuning settings across 11 downstream tasks grouped into structural abnormality detection, atherosclerotic plaque assessment, and interventional decision prediction, using both an internal validation cohort and an independent external test cohort. RESULTS CAG-MIND demonstrated robust performance across all three task categories. In the zero-shot setting, the model achieved mean AUROCs of 0.686 in the internal validation cohort and 0.745 in the external test cohort, indicating transferable multimodal representations without task-specific supervision. Following supervised fine-tuning, the mean AUROC increased to 0.827 and 0.846, respectively, with excellent performance for coronary stenosis detection (AUROC 0.940 in both cohorts), balloon/stent prediction (0.900 and 0.907), and CABG recommendation (0.877 and 0.875). Compared with representative biomedical vision-language models and conventional image-based architectures, CAG-MIND consistently achieved superior performance in both zero-shot and supervised settings and remained superior to fully fine-tuned competing models when trained with only 10% of the labelled data. Grad-CAM visualization demonstrated anatomically plausible lesion-focused attention, supporting the interpretability of the learned representations. CONCLUSIONS CAG-MIND is, to our knowledge, the first large-scale vision-language foundation model for coronary angiography trained at more than 100,000-patient scale. By aligning standardized multi-view angiographic videos with report-derived procedural semantics, CAG-MIND enables robust zero-shot transfer, data-efficient fine-tuning and cross-center generalization. These findings support domain-aligned multimodal pretraining as a scalable foundation-model paradigm for invasive cardiovascular imaging and future cath-lab decision support.

4
Cine Cardiac MRI Captures Cardiovascular Disease Risk Beyond Established Clinical Risk Factors: Evidence from the UK Biobank

Hasny, M.; Daza, L.; Bressem, K.; Di Folco, M.; Schnabel, J. A.

2026-08-11 health informatics 10.64898/2026.08.10.26360086 medRxiv
Top 0.1%
11.9%
Show abstract

Early and accurate risk stratification of cardiovascular disease (CVD) is crucial to initiate timely preventive interventions. As large-scale multimodal clinical cohorts become increasingly available, there is growing interest in whether incorporating additional sources of information can improve CVD risk stratification. Cine cardiac MR (CMR) represents a compelling example of such a source, as it captures objective, high-dimensional structural and functional information about the heart, independent of patient-reported data. In this study, we deploy a flexible vision-tabular method to incorporate cine CMR into CVD risk assessment together with structured clinical data. Using a large prospective imaging cohort from the UK Biobank, we show that cine CMR encodes CVD risk beyond established risk scores, increasing AUROC by 0.036 over SCORE2, the best-performing traditional risk score (0.742 vs. 0.706, \textit{p} = 0.04). Furthermore, we find that cine CMR achieves risk discrimination capabilities on par with automated, image-derived phenotypes, removing the dependency on segmentation pipelines. Lastly, we demonstrate that integrating cine CMR with clinical variables through a vision-tabular learning framework stabilizes risk prediction under real-world conditions of incomplete tabular data, a common challenge in clinical practice. Together, these findings position cine CMR as a promising modality for CVD risk assessment.

5
An Explainable and Comparative Transfer Learning Framework for Brain Tumor Classification from MRI Images

Bethala, S.; Vanshika,

2026-08-10 radiology and imaging 10.64898/2026.08.06.26359900 medRxiv
Top 0.1%
10.8%
Show abstract

Automated detection of brain tumors from Magnetic Resonance Imaging (MRI) can accelerate diagnosis and reduce inter-reader variability, yet many existing studies report only top-line accuracy on small datasets, omit efficiency analysis, and provide no interpretability, limiting their clinical credibility. We present a reproducible, comparative, and explainable transfer- learning framework for binary brain-tumor classification. Our framework (i) standardizes a configurable preprocessing pipeline combining CLAHE contrast enhancement and unsharp-mask sharpening, (ii) evaluates a custom CNN baseline and pretrained backbones under an identical training budget, (iii) reports a full metric suite (accuracy, precision, recall, F1, ROC-AUC, PR-AUC, parameter count, and inference latency), and (iv) applies Grad- CAM for spatial interpretability. On a public 253-image MRI dataset (38-image held-out test set), MobileNetV2 achieves the best overall performance (94.74% accuracy, 0.994 ROC-AUC, 0.996 PR-AUC) with only 2.59M parameters and 5.9 ms per- image inference, making it the most deployment-friendly model. Larger backbones (Xception, EfficientNetB0) and the custom CNN converge to degenerate all-positive predictions under the same limited budget, illustrating the small-data overfitting risk that accuracy-only reporting conceals. Grad-CAM confirms that the best model attends to the tumor region. All source code, con- figuration files, and trained evaluation scripts are publicly avail- able at https://github.com/blck-iris/explainable-brain-tumor-mr

6
A Vision-Language Framework for Predicting Brain Tumor Recurrence from Multimodal, Longitudinal Patient Data

Tak, D.; Sreedhar, D.; Aerts, H.; Kann, B.

2026-08-12 pediatrics 10.64898/2026.08.11.26360196 medRxiv
Top 0.1%
7.9%
Show abstract

Accurate prediction of tumor recurrence in brain tumor patients following surgery is essential for optimizing adjuvant therapy, response assessment, and surveillance regimen. While MRI remains the gold standard for surveillance, integrating patient-specific clinical context may inform recurrence prediction. Traditional multimodal deep learning approaches often incorporate clinical data via simple fusion, failing to fully capture the semantic interdependencies between visual features and clinical context. Trained on over 5,000 scans from approximately 400 pediatric low-grade glioma subjects and validated across three institutional cohorts, including one clinical trial cohort, our experiments demonstrate incremental performance gains when progressing from vision-only to clinical-vision to a vision-language approach. Our results indicate that converting structured clinical covariates into natural language text allows for more effective synthesis of multimodal data, while providing a platform for incremental addition of clinical context without extending model complexity. We demonstrate that our proposed VLM architecture offers a promising direction for neuro-oncological prognosis by effectively encoding imaging cues and clinical context, with potential applicability to other longitudinal prognosis tasks.

7
MaternaAI: Enhancing Equitable Maternal Healthcare in Kerala with Fairness-Aware and Explainable Learning Models

Jo, A. A.

2026-08-14 obstetrics and gynecology 10.64898/2026.08.12.26360340 medRxiv
Top 0.1%
7.8%
Show abstract

Maternal healthcare prediction systems often suffer from algorithmic biases due to socio-economic disparities and imbalanced datasets, limiting their effectiveness for equitable healthcare policymaking. This paper introduces MaternaAI, a fairness-aware and explainable learning framework designed to enhance maternal healthcare predictions in Kerala, India. The framework focuses on three critical health indicators:(1) Tetanus Toxoid (TT) booster uptake,(2) immunization coverage rates, and (3) the percentage of pregnant women completing four or more Antenatal Care (ANC) visits. To address fairness, we propose Adaptive Equity Score Optimization (AESO), a novel optimization algorithm that dynamically integrates fairness constraints into model training. AESO is model-agnostic and adapts group equity weights in response to real-time disparities. We integrate SHAP, LIME, and feature permutation techniques for explainability, enabling transparent global and local interpretation. Empirical results demonstrate that MaternaAI significantly improves fairness metrics and model accuracy across diverse machine learning and deep learning models, offering interpretable and equitable decision support for public health stakeholders.

8
ClinSeg: Robust Brain Segmentation for Clinically Acquired Pediatric MRI

Levitis, E.; Tregidgo, H. F. J.; Zimmerman, D.; Jung, B.; Karandikar, S.; Gardner, M.; Mattisson, P.; Kafadar, E.; Zapaishchykova, A.; Kann, B. H.; Sotardi, S. T.; Vossough, A.; Huang, H.; Billot, B.; Iglesias Gonzales, J. E.; Alexander, D. C.; Alexander-Bloch, A. F.; Seidlitz, J.

2026-09-02 pediatrics 10.64898/2026.08.28.26361643 medRxiv
Top 0.1%
7.8%
Show abstract

Clinical brain MRIs from pediatric health systems represent a viable resource for modeling early neurodevelopmental trajectories and studying neurodevelopmental risk in real-world populations. However, a limitation to date has been the performance of existing segmentation tools for measuring various brain phenotypes in clinical scans. In particular, many tools underperform in infant scans due to morphological and physical changes such as rapid myelination. Here, we introduce ClinSeg: a robust segmentation approach tailored to early-life clinical MRIs with variable orientation, resolution, and contrast. We leverage existing registration and synthetic data generation tools to construct a training corpus for a 3d U-Net spanning anatomical and contrast diversity, including scans with morphological abnormalities from a pediatric hospital. Validated against manual segmentations, ClinSeg outperforms existing models in infancy while matching them in childhood and adolescence. Finally, ClinSeg enables the construction of reference brain growth trajectories in 11,699 individuals from 0-21 years of age, leading to the detection of more nuanced age-related findings in clinical groups.

9
LDCT-to-SDCT as a Bridge Problem: Single-Step Residual Endpoint Flow Matching for Real-Time Denoising

dela Sotta, T.; Saavedra, J. M.; Chang, V.; Xavier, A.; Henriquez, H.; Orellana, Y.; Curimil, J.

2026-08-31 radiology and imaging 10.64898/2026.08.27.26361520 medRxiv
Top 0.1%
6.7%
Show abstract

Diffusion models achieve high reconstruction quality in low-dose computed tomography (LDCT), but their iterative sampling trajectories impose substantial computational costs. Unlike unconditional generation, paired LDCT reconstruction starts from an image that already contains the anatomy and spatial structure of the standard-dose CT (SDCT) target; reconstruction primarily requires correcting dose-related noise and artifacts. We therefore introduce Residual Endpoint Flow Matching (REFM), an LDCT reconstruction method that learns to transport an LDCT image directly toward its paired SDCT endpoint rather than defining a noise-to-image trajectory. REFM predicts the residual velocity along linear interpolations between both images and supports single-step and multi-step reconstruction using the same trained network. We evaluate five model capacities using 1 to 50 Euler steps against deterministic U-Net and diffusion-based baselines. Across all REFM variants, one-step inference consistently provides the highest reconstruction quality. On the TCIA validation set, REFM Base achieves 50.98 dB PSNR and 0.9865 SSIM at 94.54 fps, compared with 50.92 dB, 0.9847, and 9.26 fps for DDPM-10. REFM Small retains 50.71 dB while increasing throughput to 198.56 fps. Without fine-tuning, REFM Base also matches the 25-step DDPM baseline on the external Mayo Clinic dataset, although DDPM remains stronger on synthetically degraded CRLM images. Thus, our results show that exploiting paired anatomical correspondence enables diffusion-level LDCT reconstruction with a single step reconstruction.

10
Clinically Generalisable End-to-End Graph Learning for CT Image-Based Multitask Stroke Diagnosis

Lu, Z.; Uddin, S.; Uribe, S.; White, S.; Martins, R. T.; Chau, S.; Mosaddek, A. S. M.; Islam, M. S.; Nahar, N.; Azad, A. K. M.; Hossain, K. M. N.; Choudhury, H. S.; Hasan, K. M. R.; Mosaddek, N.; Rahman, S.; Hossain, M. M.; Sizar, K. M. M. H.; Angione, C.; Lio, P.; Islam, M. T.; Moni, M. A.

2026-08-31 radiology and imaging 10.64898/2026.08.26.26360026 medRxiv
Top 0.1%
6.3%
Show abstract

Stroke remains a leading cause of mortality and long-term disability worldwide, yet rapid diagnosis is often limited by the shortage of trained radiologists, particularly in resource-constrained settings. Automated analysis of CT imaging offers a potential solution, but existing methods often struggle to achieve clinically generalisable performance while jointly addressing multiple diagnostic tasks. Here we present the Intelligent Integrated Stroke Diagnosis System IISDS, an end-to-end deep learning framework built upon StrokeGNN, a graph-based architecture that integrates 3D contextual feature extraction with U-Net-based 2D lesion segmentation to enable comprehensive stroke analysis from non-contrast CT scans. IISDS performs stroke subtype classification, lesion segmentation and lesion volume estimation within a unified pipeline. To develop and validate the system, we collected and curated BGD-ISD through a collaboration between AI researchers, neurologists, radiologists and clinicians, resulting in a large multi-centre dataset comprising 1,507 CT scans from 597 stroke cases acquired across six hospitals and medical centres in Bangladesh. Across BGD-ISD and multiple publicly available datasets, IISDS achieves state-of-the-art performance on all tasks, improving segmentation accuracy by [≥]0.011 Dice score, reducing lesion volume estimation error by [≥]0.3 average symmetric surface distance (ASSD), and increasing classification performance by [≥]0.018 area under the receiver operating characteristic curve (AUC) compared with existing approaches. These results demonstrate the potential of graph-based deep learning to enable clinically generalisable, automated and scalable stroke diagnosis from CT imaging, supporting rapid clinical decision-making, particularly in healthcare environments with limited access to expert radiological interpretation.

11
Whole-brain modeling of dynamic causal circuits in human cognition using amortized variational inference

Lee, B.; Rouillard, L.; Diniz, L. L.; Jiang, L.; Ambrogioni, L.; Ryali, S.; Branigan, N.; Mistry, P.; Cai, W.; Wassermann, D.; Menon, V.

2026-08-09 neuroscience 10.64898/2026.08.03.742253 medRxiv
Top 0.1%
6.2%
Show abstract

Understanding dynamic mechanisms underlying cognition remains a major challenge in human neuroscience. Here, we develop, validate, and apply Multivariate Dynamical Systems Identification with Amortized Variational Inference (MDSI-AVI), a novel computational framework designed to address critical challenges in capturing asymmetric, context-dependent, whole-brain directed interactions while accounting for regional hemodynamic response variability in fMRI data. MDSI-AVI leverages simulation-based inference through forward and reverse variational inference to address the limitations of conventional variational methods in high-dimensional settings. By averaging over uncertainty in hemodynamic response parameters using forward simulation, MDSI-AVI provides well-calibrated posteriors of directed connectivity that scale efficiently to networks with hundreds of nodes. Applied to Human Connectome Project data (N=728), MDSI-AVI reveals new insights into working memory mechanisms, identifying the dorsal anterior insula as a critical hub influencing activity at the whole-brain level. We demonstrate task-dependent modulation of causal influences, where the salience network drives frontoparietal network activity, which differentially influences the default mode and sensorimotor networks depending on working memory load. These whole-brain causal interactions distinguish task conditions with high accuracy and predict working memory performance. Our framework demonstrates reproducible results across whole-brain parcellations, establishing MDSI-AVI as a robust tool for advancing our understanding of circuit dynamics in cognition and disease.

12
SILICA: Streamline Independent Component Analysis for Trajectory-Resolved White Matter Decomposition

Wu, L.; Calhoun, V.

2026-08-12 neuroscience 10.64898/2026.08.06.743368 medRxiv
Top 0.1%
6.2%
Show abstract

Whole-brain tractography reconstructs the major white matter pathways as millions of individual streamlines, offering an exceptionally rich description of neural geometry. Yet the statistical methods used to compare these reconstructions across individuals inevitably discard key information. Voxel-based analyses sacrifice pathway continuity, trajectory-based methods rarely support population-level statistical decomposition, and connectome models largely abstract away the underlying geometry. No existing framework jointly characterizes the population-level statistical organization of white matter and the three-dimensional geometry of the pathways from which that organization is expressed. We introduce streamline independent component analysis (SILICA), a framework that links group-level voxel-space statistical decomposition to subject-specific trajectories through a sparse streamline-by-voxel fingerprint. Each streamline is represented by its physical path length within a common anatomical voxel grid while retaining an explicit index-level link to its original trajectory. A two-stage dimensionality reduction reconciles tractograms of differing size and enables continuous component loadings to be back-reconstructed for every original streamline. These subject-specific loadings support weighted trajectory visualization and can be projected into voxel space to generate track-weighted component maps for conventional image-based visualization and future voxel-wise analysis. Separately, the learned group spatial components can be expressed on an independently reconstructed representative whole-brain tractogram to generate a compact trajectory-resolved atlas for group-level visualization. SILICA is a single decomposition expressed simultaneously in statistical and geometric form. SILICA was evaluated in diffusion MRI tractograms from 30 healthy adults. The recovered spatial patterns correspond to recognizable commissural, projection, and association systems. Back-reconstructions preserved individual trajectory variation while isolating components shared across the group, and their projection into voxel and trajectory space yielded interpretable maps and atlases. As a proof of concept, SILICA has not yet been validated against anatomical reference standards or evaluated for reproducibility and performance relative to established methods. Nevertheless, these results establish a coherent foundation for analyzing white matter in a framework that jointly represents population-level statistical structure and streamline geometry.

13
Benchmarking the robustness of segmentation models to corruptions in biological imaging

Kesenci, Y.; Le Folgoc, L.; Angelini, E.

2026-08-25 bioinformatics 10.64898/2026.08.21.746302 medRxiv
Top 0.1%
5.6%
Show abstract

Deep-learning-based segmentation algorithms have gained considerable accuracy for processing biological images. In particular, the introduction of large foundation models, novel architectures, and semantically varied datasets now allows for deployment of state-of-the-art models for clean image cohorts with limited re-training or, in the best of cases, in an out-of-the-box fashion. Biological imaging, however, is liable to corruptions that can hinder their deployment. While some methods document their robustness to the most common corruptions, a systematic robustness analysis of the state of the art to the expansive gamut of corruptions in biological imaging remains to be done. We perform this benchmarking by simulating 36 corruption types with varying degradation severity on images sampled from 30 different datasets. Our benchmark accounts both for the variety in biological images and the nature of corruptions. Among other things, our study reveals that performance on clean images does not correlate with overall robustness to image corruptions. In fact, we find that a decade-old method, StarDist, is more robust than many of its more recent foundation-model-based counterparts. We also show in a dedicated representation analysis that the performance of segmentation models collapses in the early layers of the encoding phase.

14
A Time-Dependent Diffusion MRI Framework for Clinical Characterisation of Human Brain Cellular Architecture

Leibovici, A.; Espinos Soler, E.; Mesika, D.; Tsarfaty, G.; Livny, A.; De Santis, S.; Eggl, M. F.

2026-09-05 radiology and imaging 10.64898/2026.09.02.26362017 medRxiv
Top 0.1%
5.5%
Show abstract

Diffusion-weighted MRI, beyond the commonly used diffusion tensor framework, offers a unique window into tissue microstructure in vivo, yet its clinical adoption has remained limited. Major barriers include the complexity of diffusion MRI sequence design, lengthy acquisition protocols, and the challenges associated with robust estimation of high-dimensional microstructural model parameters. Here, we address these limitations by combining optimised diffusion encoding with state-of-the-art simulation-based inference, establishing a clinically feasible framework for multi-compartment diffusion modelling. We validate the approach through i) in-depth in silico experiments and ii) in vivo studies made up of both human and rodent data. The resulting microstructural metrics are robust, reproducible across healthy individuals and show significant spatial associations with brain-wide expression patterns of cell-specific genes. Requiring less than 10 minutes of acquisition time, this framework substantially lowers the barriers to advanced microstructural imaging, a prerequisite step toward its eventual evaluation for the diagnosis, stratification, and monitoring of brain disorders.

15
A generative model for dimensionality reduction with millions of features and few samples

Pancotti, C.; Fariselli, P.; Meisner, J.; Krogh, A.

2026-08-09 bioinformatics 10.64898/2026.08.04.742788 medRxiv
Top 0.2%
4.4%
Show abstract

MotivationIn this paper, we demonstrate that it is feasible to train a deep generative model for dimensionality reduction with millions of features using few samples, which makes this type of generative model a more versatile alternative to standard methods for dimensionality reduction. Specifically, we hypothesize that for a decoder-only model, the number of training samples required is almost independent of the feature dimensionality in most network architectures. ResultsThrough an extensive set of experiments on synthetic non-linear data, we validate this hypothesis. We also train the model on a downsampled version of the 1000 Genomes Project (1KGP) dataset to further assess its behavior under controlled reductions in sample size. Furthermore, we train a deep generative decoder (DGD) on a curated dataset from the International Cancer Genome Consortium (ICGC), which contains 4.4 million features. It is trained on approximately 4,000 samples and tested on 1,000 samples. The resulting latent representation exhibits clear clustering, and when methods are reduced to the same number of dimensions, it outperforms PCA and VAE for tumor type classification. Additionally, the DGD is computationally efficient and can be trained on a 16GB GPU. Availability and implementationCode is available at https://github.com/cpancott/ReceptiveDGD. Contactcorrado.pancotti@helmholtz-munich.de; akrogh@di.ku.dk Supplementary informationSupplementary data are available with this preprint.

16
Augmenting Deep Learning-Based PSMA PET/CT Metastasis Segmentation with a Population-Level Spatial Atlas

Chau, G. N.; Biswas, B. A.; Wagle, B. R.; Maeder, M. E.; Yu, J. B.; Bhattacharya, I.

2026-08-31 radiology and imaging 10.64898/2026.08.26.26361439 medRxiv
Top 0.2%
4.3%
Show abstract

Automated lesion segmentation is increasingly central to PSMA PET/CT interpretation, supporting staging, treatment planning, and response assessment at a scale that outpaces available nuclear-medicine expertise. However, automated PSMA-PET/CT whole-body lesion segmentation models are trained on images alone, with no knowledge of where in the body prostate metastases actually tend to occur. Radiologists use clinical domain knowledge of metastatic spread, but its absence in machine learning models produces false positives in anatomically implausible locations and missed lesions in high-risk sites such as the liver. In this work, we explore whether population-level spatial knowledge of metastatic spread can be used to augment deep learning segmentation predictions, and how such a prior should be fused with a network's output, without additional training. We build a data-driven metastasis atlas from 375 expert-annotated whole-body PSMA PET/CT scans and investigate its fusion with a trained segmentation network under a Bayesian framework, in which prediction probabilities from an nnU-Net-based lesion segmentation model serve as the likelihood and the data-driven atlas as the prior. Because metastases occupy only a small fraction of whole-body voxels, the atlas's peak probability is too low, and standard power-scaled or naive Bayesian pooling references lack the tools to deal with this shortcoming. This causes these standard fusion strategies to fail and, in the naive Bayesian case, to sharply degrade performance. We instead derive a calibrated, background-referenced log-odds fusion, one of many possible approaches to combine a population atlas with a deep learning model's predictions, distinct from classical multi-atlas label fusion in that it fuses a single population prior with a trained network's softmax rather than combining several registered atlases. Furthermore, this approach is neutral outside atlas support by construction, reduces exactly to the baseline network when unweighted, and requires no retraining. This atlas fusion significantly improved mean Dice over the baseline nnU-Net on a disjoint internal test set ($+0.011$, Holm-adjusted $p=0.021$) and on an independent external cohort ($+0.0129$, Holm-adjusted $p=3.8\times10^{-16}$), with lesion sensitivity improving from 0.849 to 0.861 internally and Dice improving over baseline in every stratified anatomic region, including the rare, high-risk sites motivating this work, while naive Bayesian pooling degrades performance sharply and power-scaled pooling underperforms it throughout. Our findings suggest that population-level spatial priors can meaningfully augment deep learning predictions in whole-body oncologic segmentation, provided the fusion rule is calibrated to where the prior actually carries signal.

17
Construction of a Standardized Time-Lapse Imaging Database and a Gradient Boosting Ensemble Framework for Integrating Zygote Morphokinetic Parameters with Conventional Embryo Assessment

ZHAO, M.; LIU, J.; HAN, D.; ZHANG, C.; ZHOU, Y.; CHEN, S.; LIU, C.

2026-08-24 obstetrics and gynecology 10.64898/2026.08.20.26359523 medRxiv
Top 0.2%
4.1%
Show abstract

In vitro fertilization (IVF) laboratories equipped with timelapse incubators generate vast quantities of sequential embryo images, yet the absence of standardized, annotated databases impedes the development of reproducible computational tools for embryo assessment. Here we describe the construction of a standardized time-lapse imaging database comprising 631 normally fertilized zygotes from 218 treatment cycles, integrating timelapse image sequences, patient clinical records, and embryo developmental outcomes. We further present a gradient boosting decision tree (GBDT) ensemble framework that integrates zygote morphokinetic parameters-continuous time-series features extracted via a validated CNN-based segmentation algorithm (US Patent US11210494B2)-with conventional embryo assessment grades (categorical features per the Istanbul consensus). The fusion framework employs equal-weight initialization followed by iterative residual-decreasing training to optimally combine heterogeneous feature types. Ablation analysis demonstrated that the integrated model achieved an AUC of 0.78, significantly outperforming morphokinetics-only (AUC 0.71) and conventional-only (AUC 0.65) models, confirming the complementary value of the two data modalities. The database and fusion framework provide a reproducible foundation for embryo development assessment and are generalizable to other multimodal data integration tasks in reproductive medicine.

18
ALFIE: Anatomy-aware enhancement of Low FIEld 64mT T2-weighted neonatal brain MRI for structural analysis

Cawley, P.; Uus, A.; Colford, K.; Padormo, F.; Teixeira, R.; Tomazinho, I.; UNITY Consortium, ; Williams, S. C. R.; Edwards, A. D.; O'Muircheartaigh, J.; Arichi, T.; Hajnal, J. V.; Rutherford, M. A.

2026-08-28 pediatrics 10.64898/2026.08.25.26361317 medRxiv
Top 0.2%
4.0%
Show abstract

Purpose: To develop and evaluate an anatomy-aware deep learning framework for enhancement of neonatal 64mT T2-weighted MRI that improves anatomical visibility while preserving native ultra-low-field contrast and enabling quantitative structural analysis. Methods: A multitask network, jointly performing image enhancement and tissue segmentation, was trained on 75 and evaluated on 20 paired neonatal 64mT/3T MRI datasets spanning a broad range of gestational ages and pathologies. To preserve native 64mT contrast, 3T images were locally harmonized before training. The framework also generated quality-control maps and regional volumetric measurements. Volumetric agreement was further assessed in 40 paired term-born control datasets. Results: Enhanced 64mT images showed improved image quality metrics and better delineation of cortical, deep gray matter, ventricular, white matter, and posterior fossa structures while maintaining native contrast characteristics. Tissue segmentations demonstrated good agreement with reference 3T labels. Volumetric measurements showed excellent correspondence with 3T across major tissue compartments, with only small systematic regional biases. Conclusions: Anatomy-aware enhancement enables automated tissue segmentation and volumetric analysis directly from neonatal 64mT MRI while preserving native image contrast. These findings support the feasibility of quantitative neonatal neuroimaging at ultra-low field.

19
Tractography from Serial Optical Coherence Tomography: How and Why?

Poirier, C.; Petit, L.; Lefebvre, J.; Descoteaux, M.

2026-08-19 bioinformatics 10.64898/2026.08.14.744847 medRxiv
Top 0.2%
3.2%
Show abstract

To disentangle complex fiber configurations that remain challenging for diffusion MRI tractography, insights might be gained from microscopy tractography. Indeed, by precisely following small white matter (WM) fascicles invisible at the resolution of diffusion MRI, microscopy tractography can help explain how fiber populations are organized at the finest scales. Serial optical coherence tomography (S-OCT) is an imaging modality relying on the intrinsic contrast of a sample. When applied to brain tissues, the S-OCT contrast is primarily driven by the myelin reflectivity. Due to its high resolution, on the order of microns, and its 3D nature, S-OCT offers promise for studying WM connections at the microscale. However, while other microscopy imaging modalities have been shown to enable tractography, whether the reflectivity contrast from S-OCT supports the reconstruction of long-range WM fascicles at the microscale remains unknown. Furthermore, there is a gap in the literature regarding how an ideal microscopy tractography algorithm should behave with respect to the choice of tractography algorithm, tracking maps definition and microscale orientation distribution functions (ODF) estimation. In this work, we describe a tailored approach to reconstruct WM fascicles at the microscale from S-OCT acquisitions. We improve microscale orientation distribution functions (ODF) estimation by implementing a sliding-window formulation allowing the estimation of ODF at S-OCT resolution, and use apodized Dirac delta functions for reducing unwanted interference. We validate our approach on a simulated microscopy-like FiberCup dataset, and show that using multiscale Frangi filters for estimating ODF outperforms structure tensor analysis. We also show that particle filtering tractography with anatomical constraints enables targetted, region-to-region tractography, and outperforms standard deterministic or probabilistic tracking approaches. We further demonstrate our method on a whole mouse brain S-OCT reconstruction at 10 m by reconstructing the thalamocortical white-matter projections. Overall, our results show that S-OCT tractography recovers fine white matter fascicles visible at the microscale, and that these connections are supported by viral tracing experiments from the Allen Mouse Brain Connectivity Atlas. Moreover, this work shows the first ODF estimation and fully-3D probabilistic particle filtering tractography of the mouse brain from S-OCT reconstructions at 10 m isotropic resolution.

20
General-purpose time-series foundation models enable sample-efficient transfer learning in retinal electrophysiology

Porter, H. L.; Giles, C. B.; Kottapalli, S.; Wren, J. D.

2026-08-19 ophthalmology 10.64898/2026.08.17.26360646 medRxiv
Top 0.3%
2.6%
Show abstract

Electroretinography (ERG) measures the functional response of distinct retinal cells to light, but was largely displaced by structural imaging in the 2000s. Standardization efforts by the International Society for Clinical Electrophysiology of Vision (ISCEV) began in the late 1980s, and collapsed the rich time-series traces into reproducible components and implicit times. Recent improvements in hardware (RETeval) and software (artificial intelligence) may increase the utility of ERG data. However, no ERG-specific foundation models exist, and there are not enough public datasets to train one. We asked whether time-series foundation models (FMs) trained without ERG-specific pre-training could be adapted through transfer learning. Using two public datasets, PERG-IOBA (pattern ERG with ocular diagnoses), and LEOPs (full-field ERG focusing on Autism Spectrum Disorder, ASD), we interrogated how FMs could improve over smaller within-domain models. We measured the binary (healthy/typically developing vs any annotation) and multiclass (specific family/diagnosis) classification performance of both frozen and fine-tuned FMs, alongside custom autoencoder and multiscale models, using patient-aware splits for cross validation. We benchmark the same architectures against PTB-XL, a large 12-lead ECG corpus, as both a control for each approach and to explore scaling behavior. We show that 1) pre-trained FMs can reconstruct masked traces from all three datasets, 2) frozen and fine-tuned embeddings, especially combined with multimodal metadata through masked autoencoders, performed best on classification tasks. Performance on PTB-XL was maintained down to 300 records, comparable in size to the ERG datasets. We could not reproduce published classification performance on the ASD task. Taken together, these results support general purpose foundation models as a practical approach to ERG analysis.